Papers with social reasoning
Social Intelligence in the Age of LLMs (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a powerful tool for integrating human-like communication and context-aware interactions into artificial systems. |
| Approach: | They propose to introduce and overview different aspects of artificial social intelligence and their relationship with LLMs by introducing scientific methods for evaluating social intelligence in LLM. |
| Outcome: | This tutorial will introduce scientific methods for evaluating social intelligence in LLMs, highlighting the key challenges, and identifying promising research directions. |
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It (2026.findings-eacl)
Copied to clipboard
| Challenge: | Using social reasoning benchmarks, we uncover pervasive flaws in both benchmark items and evaluation methodology. |
| Approach: | They audit three widely used social reasoning benchmarks and identify flaws in their design and evaluation methodology. |
| Outcome: | The results challenge the validity of current benchmark-based claims about social reasoning in large language models. |
A Notion of Complexity for Theory of Mind via Discrete World Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Theory of Mind (ToM) can be used to assess the capabilities of Large Language Models (LLMs) in complex scenarios where social reasoning is required. |
| Approach: | They propose a framework inspired by cognitive load theory to measure the complexity of ToM tasks by a prompting technique that augments the information available to a model with a description of how the environment changes with the agents’ interactions. |
| Outcome: | The proposed framework assesses the complexity of five widely adopted ToM benchmarks and shows that it performs better than other frameworks. |
CRoW: Benchmarking Commonsense Reasoning in Real-World Tasks (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent efforts in natural language processing (NLP) commonsense reasoning research have produced a number of new datasets and benchmarks. |
| Approach: | They propose a manually-curated, multi-task benchmark that evaluates models' ability to apply commonsense reasoning in the context of six real-world NLP tasks. |
| Outcome: | The proposed benchmark evaluates the ability of models to apply commonsense reasoning in the context of six real-world NLP tasks. |
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs (2026.acl-long)
Copied to clipboard
| Challenge: | Recent studies show that large language models fail to challenge users’ harmful beliefs in domains ranging from medical advice to social reasoning. |
| Approach: | They propose to examine whether pragmatic factors influence LLM accommodation and epistemic vigilance in humans. |
| Outcome: | The proposed model can be understood and addressed as having excessive accommodation and insufficient epistemic vigilance. |
Are they lovers or friends? Evaluating LLMs’ Social Reasoning in English and Korean Dialogues (2026.acl-long)
Copied to clipboard
Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Doğruöz, Alice Oh, Najoung Kim
| Challenge: | Existing studies on LLMs' ability to infer social relationships have limited results for Korean and English. |
| Approach: | They propose a social reasoning task based on a 1.1k-dialogue dataset in English and Korean sourced from movie scripts to evaluate LLMs' ability to infer the social relationships between speakers. |
| Outcome: | The proposed task evaluates the ability of LLMs to infer the social relationships between speakers in 1.1k-dialogue datasets in English and Korean. |
VIBE: Can a VLM Read the Room? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Vision Language Models (LLMs) cannot account for the role that non-verbal cues play in understanding social situations. |
| Approach: | They propose a task to test the capabilities of Vision Language Models (VLMs) to account for the visual social-pragmatic inference gap. |
| Outcome: | The proposed task tests the capabilities of a VLM for a social reasoning task. |
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Social reasoning is a core competency of social intelligence and requires specialized neural and cognitive systems to be able to interpret multimodal interactions. |
| Approach: | They propose to use social reasoning traces to generate fine-grained explanations using external knowledge. |
| Outcome: | The proposed model is based on 272 videos of human interactions and 1,486 human-annotated reasoning traces related to inferences about these interactions. |
Bayesian Social Deduction with Graph-Informed Language Models (2026.acl-long)
Copied to clipboard
Shahab Rahimirad, Guven Gergerli, Lucia Romero, Angela Qian, Matthew Lyle Olson, Simon Stepputtis, Joseph Campbell
| Challenge: | Large language models (LLMs) have demonstrated remarkable general-purpose reasoning capabilities across a wide range of tasks. |
| Approach: | They propose a hybrid reasoning framework that externalizes belief inference to a structured probabilistic model while using an LLM for language understanding and interaction. |
| Outcome: | The proposed framework achieves competitive performance with larger models in Agent-Agent play and is the first language agent to defeat human players in a controlled study. |